Distral: Robust multitask reinforcement learning

نویسندگان

Yee Whye Teh

Victor Bapst

Wojciech Czarnecki

John Quan

James Kirkpatrick

Raia Hadsell

Nicolas Heess

Razvan Pascanu

چکیده

Most deep reinforcement learning algorithms are data inefficient in complex and rich environments, limiting their applicability to many scenarios. One direction for improving data efficiency is multitask learning with shared neural network parameters, where efficiency may be improved through transfer across related tasks. In practice, however, this is not usually observed, because gradients from different tasks can interfere negatively, making learning unstable and sometimes even less data efficient. Another issue is the different reward schemes between tasks, which can easily lead to one task dominating the learning of a shared model. We propose a new approach for joint training of multiple tasks, which we refer to as Distral (Distill & transfer learning). Instead of sharing parameters between the different workers, we propose to share a “distilled” policy that captures common behaviour across tasks. Each worker is trained to solve its own task while constrained to stay close to the shared policy, while the shared policy is trained by distillation to be the centroid of all task policies. Both aspects of the learning process are derived by optimizing a joint objective function. We show that our approach supports efficient transfer on complex 3D environments, outperforming several related methods. Moreover, the proposed learning process is more robust and more stable—attributes that are critical in deep reinforcement learning.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Multitask Representation Using Reusable Local Policy Templates

Constructing robust controllers to perform tasks in large, continually changing worlds is a difficult problem. A long-lived agent placed in such a world could be required to perform a variety of different tasks. For this to be possible, the agent needs to be able to abstract its experiences in a reusable way. This paper addresses the problem of online multitask decision making in such complex w...

متن کامل

Hierarchy Through Composition with Multitask LMDPs

Hierarchical architectures are critical to the scalability of reinforcement learning methods. Most current hierarchical frameworks execute actions serially, with macro-actions comprising sequences of primitive actions. We propose a novel alternative to these control hierarchies based on concurrent execution of many actions in parallel. Our scheme exploits the guaranteed concurrent compositional...

متن کامل

Reinforcement Learning Based PID Control of Wind Energy Conversion Systems

In this paper an adaptive PID controller for Wind Energy Conversion Systems (WECS) has been developed. Theadaptation technique applied to this controller is based on Reinforcement Learning (RL) theory. Nonlinearcharacteristics of wind variations as plant input, wind turbine structure and generator operational behaviordemand for high quality adaptive controller to ensure both robust stability an...

متن کامل

Scalable Multitask Policy Gradient Reinforcement Learning

Policy search reinforcement learning (RL) allows agents to learn autonomously with limited feedback. However, such methods typically require extensive experience for successful behavior due to their tabula rasa nature. Multitask RL is an approach, which aims to reduce data requirements by allowing knowledge transfer between tasks. Although successful, current multitask learning methods suffer f...

متن کامل

PAC Continuous State Online Multitask Reinforcement Learning with Identification

One key feature of a general intelligent autonomous agent is to be able to learn from past experience to improve future performance. In this paper we consider how an agent can leverage prior experience from performing reinforcement learning in order to learn faster in future tasks. We introduce the first, to our knowledge, probably approximately correct (PAC) RL algorithm COMRLI for sequential ...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره شماره

صفحات -

تاریخ انتشار 2017

Distral: Robust multitask reinforcement learning

نویسندگان

چکیده

منابع مشابه

A Multitask Representation Using Reusable Local Policy Templates

Hierarchy Through Composition with Multitask LMDPs

Reinforcement Learning Based PID Control of Wind Energy Conversion Systems

Scalable Multitask Policy Gradient Reinforcement Learning

PAC Continuous State Online Multitask Reinforcement Learning with Identification

عنوان ژورنال:

اشتراک گذاری